Visual-Tactile Multimodal Learning via Hierarchical Tactile-Guided Spatial Attention and Dynamic Temporal Fusion
FAN Chenglong1, HU Lihua1, HU Jianhua2
1. School of Computer Science and Technology, Taiyuan University of Science and Technology, Taiyuan 030024; 2. Engineering Laboratory for Intelligent Industrial Vision, Institute of Automation, Chinese Academy of Sciences, Beijing 100190
摘要 视觉与触觉在感知范围和信息形式上存在差异,现有方法难以有效关联触觉接触信息与视觉局部区域,并且固定融合策略难以适应不同交互阶段两种模态贡献的动态变化.针对上述问题,文中提出基于分层触觉引导空间注意力和动态时序融合的视觉-触觉多模态学习方法.首先,以触觉特征为查询条件,引导视觉分支关注接触关联区域,并融合中、高层视觉特征,兼顾局部纹理细节与高层语义信息.然后,采用双流Mamba对视觉序列与触觉序列进行时序建模,并通过动态门控机制自适应调节两种模态的融合比例.最后,利用时序注意力池化模块对融合序列进行加权汇聚,突出关键交互时刻并抑制冗余时间步的干扰.在Touch and Go材料识别数据集上的实验表明,文中方法的材料类别识别准确率较高,消融实验和可视化分析进一步验证方法各组成模块的有效性.
Abstract:Visual and tactile modalities differ in perceptual scope and information form. Existing methods fail to effectively associate tactile contact information with local visual regions. Moreover, fixed fusion strategies cannot adapt to the dynamic variations in the contributions of the two modalities across different interaction stages. To address these issues, a visual-tactile multimodal learning method via hierarchical tactile-guided spatial attention and dynamic temporal fusion(HTA-DTF) is proposed. First, tactile features are utilized as queries to guide the visual branch to focus on contact-related regions, while intermediate- and high-level visual features are integrated to capture both local texture details and high-level semantic information. Second, a dual-stream Mamba architecture is employed to model the temporal dependencies of visual and tactile sequences, and a dynamic gating mechanism is adopted to adaptively adjust the fusion proportions of the two modalities across different interaction stages. Finally, temporal attention pooling performs weighted aggregation over the fused sequence, emphasizing key interaction moments while suppressing interference from redundant time steps. Experiments on the Touch and Go material recognition dataset demonstrate that HTA-DTF achieves high material recognition accuracy. Ablation studies and visualization analyses further verify the effectiveness of the proposed components.
[1] Lederman S J, Klatzky R L.Haptic perception: a tutorial[J]. Atten-tion, Perception, & Psychophysics, 2009, 71(7): 1439-1459. [2] Luo S, Bimbo J, Dahiya R, et al. Robotic tactile perception of object properties: a review[J]. Mechatronics, 2017, 48: 54-67. [3] Calandra R, Owens A, Upadhyaya M, et al. The feeling of success: does touch sensing help predict grasp outcomes[C]//Proceedings of the 1st International Conference on Robot Learning. San Diego, USA: JMLR, 2017: 314-323. [4] Lee M A, Zhu Y K, Srinivasan K, et al. Making sense of vision and touch: self-supervised learning of multimodal representations for contact-rich tasks[C]//Proceedings of the IEEE International Conference on Robotics and Automation. Washington, USA: IEEE, 2019: 8943-8950. [5] Cui S W, Wang R, Wei J H, et al. Self-attention based visual-tac-tile fusion learning for predicting grasp outcomes[J]. IEEE Robotics and Automation Letters, 2020, 5(4): 5827-5834. [6] Yang F Y, Ma C Y, Zhang J Z, et al. Touch and Go: learning from human-collected vision and touch[C]//Proceedings of the 36th International Conference on Neural Information Processing Systems. Cambridge, USA: MIT Press, 2022: 8081-8103. [7] Kerr J, Huang H, Wilcox A, et al. Self-supervised visuo-tactile pre-training to locate and follow garment features[EB/OL].[2026-04-17]. https://arxiv.org/pdf/2209.13042. [8] Dave V, Lygerakis F, Rueckert E.Multimodal visual-tactile representation learning through self-supervised contrastive pre-training[C]//Proceedings of the IEEE International Conference on Robotics and Automation. Washington, USA: IEEE, 2024: 8013-8020. [9] Wu Z Y, Zhao Y Q, Luo S.ConViTac: aligning visual-tactile fusion with contrastive representations[C]//Proceedings of the IEEE/RSJ International Conference on Intelligent Robots and Systems. Wa-shington, USA: IEEE, 2025: 8545-8552. [10] Lin T Y, Dollar P, Girshick R, et al. Feature pyramid networks for object detection[C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. Washington, USA: IEEE, 2017: 936-944. [11] Xu Z Q J, Zhang Y Y, Luo T, et al. Frequency principle: Fourier analysis sheds light on deep neural networks[J]. Communications in Computational Physics, 2020, 28(5): 1747-1767. [12] Anderson P, He X D, Buehler C, et al. Bottom-up and top-down attention for image captioning and visual question answering[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Washington, USA: IEEE, 2018: 6077-6086. [13] Hochreiter S, Schmidhuber J.Long short-term memory[J]. Neural Computation, 1997, 9(8): 1735-1780. [14] Bai S J, Kolter J Z, Koltun V.An empirical evaluation of generic convolutional and recurrent networks for sequence modeling[EB/OL]. [2026-04-17].https://arxiv.org/pdf/1803.01271. [15] Vaswani A, Shazeer N, Parmar N, et al. Attention is all you need[C]//Proceedings of the 31st International Conference on Neural Information Processing Systems. Cambridge, USA: MIT Press, 2017: 6000-6010. [16] Doherty J, Gardiner B, Siddique N, ,et al. A novel visuo-tactile object recognition pipeline using transformers with feature level fusion[C/OL]//Proceedings of the International Joint Conference on Neural Networks. Washington. A novel visuo-tactile object recognition pipeline using transformers with feature level fusion[C/OL]//Proceedings of the International Joint Conference on Neural Networks. Washington, USA: IEEE, 2024. https://ieeexplore.ieee.org/stamp/stamp.jsp?tp=&arnumber=10650147. [17] Jiang C P, Xu W Q, Li Y T, ,et al. Capturing forceful interaction with deformable objects using a deep learning-powered stretchable tactile array[J/OL]. Nature Communications, 2024, 15. https://www.nature.com/articles/s41467-024-53654-y.pdf. [18] Gu A, Goel K, Re C.Efficiently modeling long sequences with structured state spaces[EB/OL]. [2026-04-17].https://arxiv.org/pdf/2111.00396v3. [19] Gu A, Dao T.Mamba: linear-time sequence modeling with selective state spaces[EB/OL]. [2026-04-17].https://arxiv.org/pdf/2312.00752. [20] Liu Y, Tian Y J, Zhao Y Z, et al. VMamba: visual state space model[C]//Proceedings of the 38th International Conference on Neural Information Processing Systems. Cambridge, USA: MIT Press, 2024: 103031-103063. [21] Hatamizadeh A, Kautz J.MambaVision: a hybrid Mamba-Transformer vision backbone[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Washington, USA: IEEE, 2025: 25261-25270. [22] Arevalo J, Solorio T, Montes-Y-Gómez M, et al. Gated multimodal networks[J]. Neural Computing and Applications, 2020, 32(14): 10209-10228. [23] Cao G Q, Zhou Y, Bollegala D, et al. Spatio-temporal attention model for tactile texture recognition[C]//Proceedings of the IEEE/RSJ International Conference on Intelligent Robots and Systems. Washington, USA: IEEE, 2020: 9896-9902. [24] Ilse M, Tomczak J, Welling M.Attention-based deep multiple instance learning[C]//Proceedings of the 35th International Confe-rence on Machine Learning. San Diego, USA: JMLR, 2018: 2127-2136. [25] GAO R H, DOU Y M, LI H, et al. The ObjectFolder benchmark: multisensory learning with neural and real objects[C]//Procee-dings of the IEEE/CVF Conference on Computer Vision and Pa-ttern Recognition. Washington, USA: IEEE, 2023: 17276-17286. [26] Deng J, Dong W, Socher R, et al. ImageNet: a large-scale hierarchical image database[C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. Washington, USA: IEEE, 2009: 248-255. [27] Loshchilov I, Hutter F.Decoupled weight decay regularization[EB/OL]. [2026-04-17].https://arxiv.org/pdf/1711.05101. [28] Srivastava N, Hinton G, Krizhevsky A, et al. Dropout: a simple way to prevent neural networks from overfitting. Journal of Machine Learning Research, 2014, 15: 1929-1958. [29] Loshchilov I, Hutter F.SGDR: stochastic gradient descent with warm restarts[EB/OL]. [2026-04-17].https://arxiv.org/pdf/1608.03983.